Papers with reasoning enhancement

2 papers
MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances have enabled MLLMs to tackle complex challenges such as mathematical reasoning and multimodal understanding.
Approach: They propose a multimodal refinement benchmark to evaluate the refinement capabilities of Multimodal Large Language Models (MLLMs) the benchmark categorizes errors into six error types to highlight areas for improvement in effective reasoning enhancement.
Outcome: The proposed framework evaluates the refinement capabilities of multimodal large language models across six scenarios.
Why Did Apple Fall: Evaluating Curiosity in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for evaluating curiosity-like behaviors in large language models lack curiosity-inspired features.
Approach: They propose a psychology-inspired framework to evaluate curiosity in large language models . they adapt the Five-Dimensional Curiosity scale Revised (5DCR) to LLMs .
Outcome: The proposed framework evaluates curiosity in large language models using questionnaires and behavioral studies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations